首页> 外文OA文献 >How is a data-driven approach better than random choice in label space division for multi-label classification?

【2h】

How is a data-driven approach better than random choice in label space division for multi-label classification?

机译：数据驱动方法如何比标签空间中的随机选择更好多标签分类的划分？

代理获取

本网站仅为用户提供外文OA文献查询和代理获取服务，本网站没有原文。下单后我们将采用程序或人工为您竭诚获取高质量的原文，但由于OA文献来源多样且变更频繁，仍可能出现获取不到、文献不完整或与标题不符等情况，如果获取不到我们将提供退款服务。请知悉。

页面导航

摘要
著录项
相似文献
相关主题

摘要

We propose using five data-driven community detection approaches from socialnetworks to partition the label space for the task of multi-labelclassification as an alternative to random partitioning into equal subsets asperformed by RAkELd: modularity-maximizing fastgreedy and leading eigenvector,infomap, walktrap and label propagation algorithms. We construct a labelco-occurence graph (both weighted an unweighted versions) based on trainingdata and perform community detection to partition the label set. We includeBinary Relevance and Label Powerset classification methods for comparison. Weuse gini-index based Decision Trees as the base classifier. We compare educatedapproaches to label space divisions against random baselines on 12 benchmarkdata sets over five evaluation measures. We show that in almost all cases seveneducated guess approaches are more likely to outperform RAkELd than otherwisein all measures, but Hamming Loss. We show that fastgreedy and walktrapcommunity detection methods on weighted label co-occurence graphs are 85-92%more likely to yield better F1 scores than random partitioning. Infomap on theunweighted label co-occurence graphs is on average 90% of the times better thanrandom paritioning in terms of Subset Accuracy and 89% when it comes to Jaccardsimilarity. Weighted fastgreedy is better on average than RAkELd when it comesto Hamming Loss.

机译：我们建议使用来自社交网络的五种数据驱动的社区检测方法来为多标签分类任务分配标签空间，以替代随机划分为由RAkELd执行的相等子集的方法：最大化模块化的快速贪婪和领先特征向量，信息图，助行器和标签传播算法。我们基于训练数据构造了一个标签共现图（均加权了未加权版本），并执行了社区检测以划分标签集。我们包括二进制相关性和标签Powerset分类方法进行比较。我们使用基于gini-index的决策树作为基础分类器。我们在5种评估方法的12个基准数据集上比较了受教育的方法和标签相对于随机基线的空间划分。我们表明，在几乎所有情况下，采用七种方法进行猜测的方法比所有其他方法都可能优于RAkELd，但汉明损失法则除外。我们显示，加权标签共现图上的快速贪婪和助步社区检测方法比随机分区产生更好的F1分数的可能性高85-92％。就子集准确度而言，未加权标签共现图上的信息图平均比随机划分好90％，而在Jaccardlikeness方面平均要好89％。在汉明损失方面，加权fastgreedy平均优于RAkELd。

著录项

作者
Szymański, Piotr; Kajdanowicz, Tomasz; Kersting, Kristian;
展开▼
作者单位

展开▼
年度 2016
总页数
原文格式 PDF
正文语种
中图分类

相似文献

外文文献
中文文献
专利

1. How Is a Data-Driven Approach Better than Random Choice in Label Space Division for Multi-Label Classification? [J] . Piotr Szymański, Tomasz Kajdanowicz, Kristian Kersting Entropy . 2016,第8期

机译：在多标签分类的标签空间划分中，数据驱动的方法如何比随机选择更好？
2. Correlated Multi-label Classification with Incomplete Label Space and Class Imbalance [J] . Braytee Ali, Liu Wei, Anaissi Ali, ACM transactions on intelligent systems . 2019,第5期

机译：标签空间不完整且类别不平衡的相关多标签分类
3. Deep Correlation Structure Preserved Label Space Embedding for Multi-label Classification [J] . Kaixiang Wang, Ming Yang, Wanqi Yang, JMLR: Workshop and Conference Proceedings . 2018,第1期

机译：用于多标签分类的深度相关结构保留标签空间嵌入
4. Random Forests with Random Projections of the Output Space for High Dimensional Multi-label Classification [C] . Arnaud Joly, Pierre Geurts, Louis Wehenkel European conference on machine learning and knowledge discovery in databases . 2014

机译：高维多标签分类的输出空间随机投影的随机森林
5. A Rule-Based Evolutionary Approach to Multi-Label Classification [D] . Nazmi, Shabnam. 2021

机译：基于规则的多标签分类的进化方法
6. Multi-label spacecraft electrical signal classification method based on DBN and random forest [O] . Ke Li, Nan Yu, Pengfei Li, -1

机译：基于dbn和随机森林的多标签航天器电信号分类方法
7. Random forests with random projections of the output space for high dimensional multi-label classification [O] . Joly, Arnaud, Geurts, Pierre, Wehenkel, Louis 2014

机译：用于高维多标签分类的具有输出空间随机投影的随机森林

How is a data-driven approach better than random choice in label space division for multi-label classification?

摘要

著录项

相似文献

相关主题

期刊订阅